Image-based logical document structure recognition
نویسندگان
چکیده
منابع مشابه
Knowledge-based derivation of document logical structure
The analysis of a document image to derive a symbolic description of its structure and contents involves using spatial domain knowledge to classify the different printed blocks (e.g., text paragraphs), group them into logical units (e.g., newspaper stories), and determine the reading order of the text blocks within each unit. These steps describe the conversion of the physical structure of a do...
متن کاملText Type Structure And Logical Document Structure
Most research on automated categorization of documents has concentrated on the assignment of one or many categories to a whole text. However, new applications, e.g. in the area of the Semantic Web, require a richer and more fine-grained annotation of documents, such as detailed thematic information about the parts of a document. Hence we investigate the automatic categorization of text segments...
متن کاملDocument Logical Structure Analysis Based on Perceptive Cycles
This paper describes a Neural Network (NN) approach for logical document structure extraction. In this NN architecture, called Transparent Neural Network (TNN), the document structure is stretched along the layers, allowing an interpretation decomposition from physical (NN input) to logical (NN output) level. The intermediate layers represent successive interpretation steps. Each neuron is appa...
متن کاملInformation Extraction from HTML Documents Based on Logical Document Structure
The World Wide Web presents the largest Internet source of information from a broad range of areas. The web documents are mostly written in the Hypertext Markup Language (HTML) that doesn’t contain any means for semantic description of the content and thus the contained information cannot be processed directly. Current approaches for the information extraction from HTML are mostly based on wrap...
متن کاملDocument image understanding: geometric and logical layout
Document Image Understanding encompasses the technology required to make paper documents equivalent to other computer exchange media like oppies, tapes, and cdroms. The physical reader of the paper document is the scanner just like the physical reader of the oppy is the oppy drive and the physical reader of the tape cartridge is the tape cartridge drive, and the physical reader of the cdrom is ...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
ژورنال
عنوان ژورنال: Pattern Analysis and Applications
سال: 2014
ISSN: 1433-7541,1433-755X
DOI: 10.1007/s10044-014-0412-8